使用Python/Beautiful Soup爬取价格时遇AttributeError错误求助
解决爬取Walmart商品页面时的AttributeError错误
错误信息
Exception has occurred: AttributeError 'NoneType' object has no attribute 'text' File "C:\Users\richard\Documents\python code\chatgptcodeaz_price_test.py", line 39, in <module> jsonData = json.loads(nextTag.text) ^^^^^^^^^^^^ AttributeError: 'NoneType' object has no attribute 'text'
问题原因
nextTag = soup.find("script",{"id":"__NEXT_DATA__"})返回None,说明页面中找不到目标script标签。大概率是Walmart的反爬机制将请求识别为机器人,返回了人机验证页面而非真实商品页面;也可能是页面结构发生了变更。
修复方案
1. 先验证请求是否被拦截
取消注释代码中的验证逻辑,运行后查看是否触发人机验证:
print(resp.text) if("Robot or human" in resp.text): print(True) else: print(False)
若输出True,说明请求被反爬拦截,需要优化请求参数。
2. 优化请求头,模拟真实浏览器
原User-Agent为旧版iPad系统,替换为更常用的桌面端浏览器标识,同时补充常见请求头字段:
headers = { "Referer": "https://www.google.com", "Connection": "Keep-Alive", "Accept-Language": "en-US,en;q=0.9", "Accept-Encoding": "gzip, deflate, br", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Upgrade-Insecure-Requests": "1" }
3. 给__NEXT_DATA__解析添加异常防护
即使请求正常,也可能存在页面结构变更的情况,必须添加异常处理避免崩溃:
nextTag = soup.find("script",{"id":"__NEXT_DATA__"}) try: if nextTag: jsonData = json.loads(nextTag.text) Detail = jsonData['props']['pageProps']['initialData']['data']['product']['shortDescription'] obj["detail"] = Detail else: obj["detail"] = None except Exception as e: print(f"解析详情失败: {str(e)}") obj["detail"] = None
4. 添加请求延迟
频繁请求会触发反爬,在请求前添加延迟模拟用户浏览间隔:
import time time.sleep(2) # 延迟2秒再发送请求
修改后的完整脚本
import requests from bs4 import BeautifulSoup import json import time headers = { "Referer": "https://www.google.com", "Connection": "Keep-Alive", "Accept-Language": "en-US,en;q=0.9", "Accept-Encoding": "gzip, deflate, br", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Upgrade-Insecure-Requests": "1" } target_url = "https://www.walmart.com/ip/Mobil-1-Advanced-Full-Synthetic-Motor-Oil-5W-30-5-qt/17034372?athbdg=L1102" time.sleep(2) resp = requests.get(target_url, headers=headers) if "Robot or human" in resp.text: print("请求被拦截,触发人机验证") else: soup = BeautifulSoup(resp.text, 'html.parser') obj = {} try: price_tag = soup.find("span", {"itemprop": "price"}) obj["price"] = price_tag.text.replace("Now ", "") if price_tag else None except: obj["price"] = None try: name_tag = soup.find("h1", {"itemprop": "name"}) obj["name"] = name_tag.text if name_tag else None except: obj["name"] = None try: rating_tag = soup.find("span", {"class": "rating-number"}) obj["rating"] = rating_tag.text.replace("(", "").replace(")", "").strip() if rating_tag else None except: obj["rating"] = None nextTag = soup.find("script", {"id": "__NEXT_DATA__"}) try: if nextTag: jsonData = json.loads(nextTag.text) Detail = jsonData['props']['pageProps']['initialData']['data']['product']['shortDescription'] obj["detail"] = Detail else: obj["detail"] = None except Exception as e: print(f"解析详情失败: {str(e)}") obj["detail"] = None print([obj])
内容的提问来源于stack exchange,提问作者headbangingcoder
相关产品推荐
相关产品推荐

