爬取Jumia电商网站时遇AttributeError: NoneType对象无get_text属性
解决Jumia爬虫AttributeError: 'NoneType' object has no attribute 'get_text'问题
问题场景
爬取Jumia尼日利亚站(jumia.com.ng)的iOS手机商品列表,遍历商品链接提取详情时触发以下错误:
Traceback (most recent call last): File "c:\Users\LP\Documents\jumia\jumia.py", line 32, in <module> name = soup.find('h1', class_='-fs20 -pts -pbxs').get_text(strip=True) AttributeError: 'NoneType' object has no attribute 'get_text'
错误原因
- 反爬拦截:Jumia的反爬机制识别到爬虫请求,返回的页面不是正常商品详情页(可能是验证码页、跳转页或空白内容),导致
soup.find找不到目标元素,返回None。 - 页面结构变更:目标元素的class名称或HTML结构已更新,原选择器失效。
- 商品页缺失元素:部分商品可能没有评论、评分或某些特征项,直接调用
get_text会报错。 - 请求头不完整:仅携带
User-Agent不足以模拟正常浏览器请求,被服务器识别为异常请求。
解决方案
1. 完善请求头
补充浏览器常用请求头,模拟真实访问行为:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.jumia.com.ng/ios-phones/', 'DNT': '1', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' }
2. 增加元素存在性判断
在调用get_text前先检查元素是否存在,避免None对象调用属性:
# 提取商品名称 name_tag = soup.find('h1', class_='-fs20 -pts -pbxs') name = name_tag.get_text(strip=True) if name_tag else '无名称' # 提取价格 amount_tag = soup.find('span', class_='-b -ltr -tal -fs24') amount = amount_tag.get_text(strip=True) if amount_tag else '无价格' # 提取评分 review_tag = soup.find('div', class_='stars _s _al') review = review_tag.get_text(strip=True) if review_tag else '无评分' # 提取评论数 rating_tag = soup.find('a', class_='-plxs _more') rating = rating_tag.get_text(strip=True) if rating_tag else '无评论'
3. 处理特征列表索引越界
特征列表可能不足6项,改用循环遍历而非固定索引取值:
features = soup.find_all('li', attrs={'style': 'box-sizing: border-box; padding: 0px; margin: 0px;'}) feature_texts = [feat.get_text(strip=True) for feat in features] # 打印所有提取到的特征 print('Key Features') for idx, feat in enumerate(feature_texts, 1): print(f"Feature {idx}: {feat}")
4. 添加异常捕获与请求延迟
- 捕获请求和解析时的异常,避免程序崩溃
- 添加请求间隔,降低被反爬拦截的概率:
import time for link in productlinks: try: r = requests.get(link, headers=headers) # 检查响应状态码,判断是否被拦截 if r.status_code != 200: print(f"请求失败,状态码:{r.status_code},链接:{link}") continue soup = BeautifulSoup(r.content, 'lxml') # 元素提取代码... time.sleep(1) except Exception as e: print(f"处理链接{link}时出错:{str(e)}") continue
5. 验证页面结构
如果修改后仍找不到元素,手动访问商品链接查看页面源码,确认目标元素的class名称是否变更,更新选择器。
修改后的完整代码
import requests from bs4 import BeautifulSoup import time baseurl = 'https://www.jumia.com' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.jumia.com.ng/ios-phones/', 'DNT': '1', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } productlinks = [] # 爬取商品列表链接 for x in range(1, 51): try: r = requests.get(f'https://www.jumia.com.ng/ios-phones/?page={x}#catalog-listing/', headers=headers) if r.status_code != 200: print(f"列表页{x}请求失败,状态码:{r.status_code}") continue soup = BeautifulSoup(r.content, 'lxml') productlist = soup.find_all('article', class_='prd _fb col c-prd') for product in productlist: link = product.find('a', href=True) if link: productlinks.append(baseurl + link['href']) time.sleep(0.5) except Exception as e: print(f"爬取列表页{x}出错:{str(e)}") continue print(f"共收集到{len(productlinks)}个商品链接") # 遍历商品链接提取详情 for link in productlinks: try: r = requests.get(link, headers=headers) if r.status_code != 200: print(f"商品页请求失败,状态码:{r.status_code},链接:{link}") time.sleep(1) continue soup = BeautifulSoup(r.content, 'lxml') # 提取商品信息,增加存在性判断 name_tag = soup.find('h1', class_='-fs20 -pts -pbxs') name = name_tag.get_text(strip=True) if name_tag else '无名称' amount_tag = soup.find('span', class_='-b -ltr -tal -fs24') amount = amount_tag.get_text(strip=True) if amount_tag else '无价格' review_tag = soup.find('div', class_='stars _s _al') review = review_tag.get_text(strip=True) if review_tag else '无评分' rating_tag = soup.find('a', class_='-plxs _more') rating = rating_tag.get_text(strip=True) if rating_tag else '无评论' features = soup.find_all('li', attrs={'style': 'box-sizing: border-box; padding: 0px; margin: 0px;'}) feature_texts = [feat.get_text(strip=True) for feat in features] # 打印信息 print(f"Name: {name}") print(f"Amount: {amount}") print(f"Review: {review}") print(f"Rating: {rating}") print('Key Features') for idx, feat in enumerate(feature_texts, 1): print(f"Feature {idx}: {feat}") print('') time.sleep(1) except Exception as e: print(f"处理链接{link}时出错:{str(e)}") time.sleep(1) continue
内容的提问来源于stack exchange,提问作者Miracle
相关产品推荐
相关产品推荐

