Python BeautifulSoup报'NoneType'无find_all属性错误咨询
BeautifulSoup 爬虫
AttributeError: 'NoneType' object has no attribute 'find_all'报错处理 报错核心原因
报错触发位置:
---> 14 content = s.find_all('p') AttributeError: 'NoneType' object has no attribute 'find_all'
触发逻辑非常明确:执行s = soup.find('div', class_='entry-content')时,当前解析的页面里不存在匹配该规则的div标签,BeautifulSoup的find()方法找不到目标元素时会固定返回None,后续对None值调用find_all()方法就会抛出属性不存在的错误。
部分URL能正常运行、部分报错,本质是不同URL返回的页面DOM结构不一致:
- 可能是请求被网站反爬拦截,返回了验证页、空白页、错误页,自然不存在要找的正文容器
- 可能是不同栏目/类型的页面,正文容器的class属性、标签类型不一样,不是所有页面都用
class="entry-content"的div承载正文 - 可能是部分页面失效返回404、权限不足返回403,页面结构和正常内容页完全不同
复现问题的原始代码
r = requests.get('https://www.marketsandmarkets.com/Market-Reports/rocket-missile-market-203298804.html/') soup = BeautifulSoup(r.content, 'html.parser') s = soup.find('div', class_='entry-content') content = s.find_all('p') print(content)
修复方案
按优先级依次处理:
- 补全请求基础配置,避免被反爬直接拦截
requests默认的请求头会被绝大多数商业网站识别为爬虫,直接返回异常页面。首先补上合法的User-Agent,同时增加请求状态校验,请求失败直接终止,不要继续解析错误响应:
import requests from bs4 import BeautifulSoup # 替换成自己浏览器的真实User-Agent即可 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } url = 'https://www.marketsandmarkets.com/Market-Reports/rocket-missile-market-203298804.html/' r = requests.get(url, headers=headers, timeout=10) # 状态码非200时直接抛错,避免解析错误页面 r.raise_for_status() soup = BeautifulSoup(r.content, 'html.parser')
- 所有
find()操作后增加空值判断,杜绝直接报错
永远不要假设目标元素一定存在,对查找结果先判空再执行后续操作:
s = soup.find('div', class_='entry-content') content = [] if s: content = s.find_all('p') print(content)
- 针对不同页面结构做兼容匹配
手动打开运行报错的URL,右键检查正文区域的实际容器标签,确认是否更换了class名、标签类型,可以写多规则匹配兼容不同页面:
# 兼容多种可能的正文容器class content = [] for class_name in ['entry-content', 'report-content', 'article-body', 'post-content']: s = soup.find('div', class_=class_name) if s: content = s.find_all('p') break
- 如果确认是反爬拦截导致返回异常页面,再针对性补充反爬处理逻辑:比如降低请求频率加延时、携带登录后的cookie、使用代理IP、处理JS验证等。
内容的提问来源于stack exchange,提问作者Amar
相关产品推荐
相关产品推荐

