You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用requests+BeautifulSoup编写的爬虫运行返回None,如何解决?

问题原因
  • 反爬策略拦截:目标网站对requests默认的请求头做了拦截,返回的不是正常的产品列表页面,导致查找元素时返回None
  • 元素定位方法错误:find()只会返回匹配到的第一个元素,你需要获取全部产品容器,应该使用find_all()方法,原代码先找单个section再调用find()找产品div,要么只能拿到第一个产品,要么匹配不到直接返回None
  • 缺少异常校验:链式调用元素查找方法时,没有判断前一步查找是否成功,只要某一步返回None,后续调用就会直接报错
解决方案

修正后的完整代码如下:

import requests
from bs4 import BeautifulSoup

# 添加请求头模拟浏览器访问,绕过基础反爬
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

url = "https://www.jab.de/tr/en/productadvancedsearch?searchTerm=&page=1"
website = requests.get(url, headers=headers)

# 先校验请求状态是否正常
if website.status_code != 200:
    print(f"请求失败,状态码:{website.status_code}")
    exit()

html = website.content
soup = BeautifulSoup(html,"html.parser")

# 先判断结果区域是否存在
result_section = soup.find("section",{"class":"results"})
if not result_section:
    print("未找到结果区域,可打印返回页内容排查反爬或页面结构变动问题")
    exit()

# 用find_all获取所有匹配的产品容器
urunListesi = result_section.find_all("div",{"class":"col-item details"})
print(f"共找到{len(urunListesi)}个产品")

for urun in urunListesi:
    # 增加异常捕获,避免单个元素结构异常导致程序崩溃
    try:
        link = urun.div.a.get("href")
        # 如果需要完整访问链接,可拼接域名:link = "https://www.jab.de" + link
        print(link)
        print("----------------------------\n")
    except AttributeError:
        print("当前产品结构异常,跳过")

核心修改点:

  • 增加了浏览器标识的请求头,绕过基础反爬拦截
  • 新增了请求状态、元素存在性的前置校验,提前定位问题
  • 将获取产品容器的find()改为find_all(),批量获取所有产品元素
  • 遍历过程增加异常捕获,兼容页面结构的局部变动

内容的提问来源于stack exchange,提问作者coder36

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 01:54:03