使用BeautifulSoup爬取网页时findAll返回空/None的问题排查
问题描述
尝试使用BeautifulSoup爬取目标产品页面,代码如下:
import requests from bs4 import BeautifulSoup url = "https://mokka-home.nl/product/bank-campania/" # Make a GET request to the URL response = requests.get(url) # Create a BeautifulSoup object soup = BeautifulSoup(response.content, 'html.parser') # Extract the product title title = soup.find('h1', {'class': 'product_title entry-title'}).text.strip() # Extract the product price price = soup.find('span', {'class': 'woocommerce-Price-amount amount'}).text.strip() # Extract all product images images = [img['src'] for img in soup.find_all('img', {'class': 'wp-post-image'})] # Print the scraped data print("Title: ", title) print("Price: ", price) print("Images: ", images)
执行时触发错误:
title = soup.find('h1', {'class': 'product_title'}).text.strip() AttributeError: 'NoneType' object has no attribute 'text'
移除.text.strip()后,输出结果为:
Title: None Price: None Images: []
问题原因
- 目标网站存在基础反爬机制,直接用
requests.get()发起请求时,服务器通过User-Agent识别出请求来自爬虫(默认UA为python-requests),返回的并非真实产品页面内容,导致BeautifulSoup找不到指定元素。 - 提取数据时直接调用
.text属性,未先判断元素是否存在,一旦soup.find()返回None,就会触发AttributeError。
解决方法
1. 添加请求头模拟浏览器访问
网站会通过User-Agent字段识别请求来源,添加浏览器UA可以绕过基础反爬:
import requests from bs4 import BeautifulSoup url = "https://mokka-home.nl/product/bank-campania/" # 模拟Chrome浏览器的请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } # 携带请求头发起请求 response = requests.get(url, headers=headers) # 先检查响应状态码,200表示请求成功 print("响应状态码:", response.status_code) soup = BeautifulSoup(response.content, 'html.parser')
2. 增加元素存在性判断,避免报错
提取数据前先判断是否找到目标元素,再调用.text或其他属性:
# 提取标题 title_elem = soup.find('h1', {'class': 'product_title entry-title'}) title = title_elem.text.strip() if title_elem else "未获取到标题" # 提取价格 price_elem = soup.find('span', {'class': 'woocommerce-Price-amount amount'}) price = price_elem.text.strip() if price_elem else "未获取到价格" # 提取图片 images = [img['src'] for img in soup.find_all('img', {'class': 'wp-post-image'})] if soup.find_all('img', {'class': 'wp-post-image'}) else [] # 打印结果 print("Title: ", title) print("Price: ", price) print("Images: ", images)
3. 验证响应内容
如果添加UA后仍无法获取数据,可以打印response.text查看返回内容:
print(response.text)
如果返回的是验证码页面或反爬提示,可能需要进一步处理(如使用代理、携带登录cookie等)。
内容的提问来源于stack exchange,提问作者Rayly Esta
相关产品推荐
相关产品推荐

