You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取网页时findAll返回空/None的问题排查

问题描述

尝试使用BeautifulSoup爬取目标产品页面,代码如下:

import requests
from bs4 import BeautifulSoup

url = "https://mokka-home.nl/product/bank-campania/"

# Make a GET request to the URL
response = requests.get(url)

# Create a BeautifulSoup object
soup = BeautifulSoup(response.content, 'html.parser')

# Extract the product title
title = soup.find('h1', {'class': 'product_title entry-title'}).text.strip()

# Extract the product price
price = soup.find('span', {'class': 'woocommerce-Price-amount amount'}).text.strip()

# Extract all product images
images = [img['src'] for img in soup.find_all('img', {'class': 'wp-post-image'})]

# Print the scraped data
print("Title: ", title)
print("Price: ", price)
print("Images: ", images)

执行时触发错误:

title = soup.find('h1', {'class': 'product_title'}).text.strip()
AttributeError: 'NoneType' object has no attribute 'text'

移除.text.strip()后,输出结果为:

Title:  None
Price:  None
Images:  []
问题原因
  • 目标网站存在基础反爬机制,直接用requests.get()发起请求时,服务器通过User-Agent识别出请求来自爬虫(默认UA为python-requests),返回的并非真实产品页面内容,导致BeautifulSoup找不到指定元素。
  • 提取数据时直接调用.text属性,未先判断元素是否存在,一旦soup.find()返回None,就会触发AttributeError。
解决方法

1. 添加请求头模拟浏览器访问

网站会通过User-Agent字段识别请求来源,添加浏览器UA可以绕过基础反爬:

import requests
from bs4 import BeautifulSoup

url = "https://mokka-home.nl/product/bank-campania/"

# 模拟Chrome浏览器的请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

# 携带请求头发起请求
response = requests.get(url, headers=headers)
# 先检查响应状态码,200表示请求成功
print("响应状态码:", response.status_code)

soup = BeautifulSoup(response.content, 'html.parser')

2. 增加元素存在性判断,避免报错

提取数据前先判断是否找到目标元素,再调用.text或其他属性:

# 提取标题
title_elem = soup.find('h1', {'class': 'product_title entry-title'})
title = title_elem.text.strip() if title_elem else "未获取到标题"

# 提取价格
price_elem = soup.find('span', {'class': 'woocommerce-Price-amount amount'})
price = price_elem.text.strip() if price_elem else "未获取到价格"

# 提取图片
images = [img['src'] for img in soup.find_all('img', {'class': 'wp-post-image'})] if soup.find_all('img', {'class': 'wp-post-image'}) else []

# 打印结果
print("Title: ", title)
print("Price: ", price)
print("Images: ", images)

3. 验证响应内容

如果添加UA后仍无法获取数据,可以打印response.text查看返回内容:

print(response.text)

如果返回的是验证码页面或反爬提示,可能需要进一步处理(如使用代理、携带登录cookie等)。


内容的提问来源于stack exchange,提问作者Rayly Esta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 13:07:37