使用BeautifulSoup的findAll()返回空结果,求解决方案
解决Mouthshut评论抓取无数据问题
核心原因
Mouthshut网站存在反爬机制,直接用requests.get()发起请求会被识别为非浏览器访问,返回的内容并非正常页面结构,导致BeautifulSoup解析不到任何有效标签。
修复方案
- 添加请求头模拟浏览器:给请求带上
User-Agent,让网站判定为正常用户访问 - 更换解析器:使用
lxml解析器(需提前安装:pip install lxml),解析稳定性优于默认的html.parser - 先验证响应有效性:请求后先检查状态码、打印响应内容,确认是否拿到正常页面
完整可运行代码
from bs4 import BeautifulSoup import requests # 模拟Chrome浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } url = 'https://www.mouthshut.com/product-reviews/berger-paints-reviews-925712965-page-4' response = requests.get(url, headers=headers) if response.status_code == 200: soup = BeautifulSoup(response.content, 'lxml') # 提取评论内容 print("===== 评论内容 =====") for div in soup.find_all('div', class_='more reviewdata'): for p in div.find_all('p'): print(p.text.strip()) # 提取评分(每个星级对应一个<i>标签,5星即5个标签) print("\n===== 评分信息 =====") rating_groups = soup.select('.rating .rated-star') print(f"共抓取到 {len(rating_groups)} 个星级标签(单条评论满星为5个)") # 提取评论标题与日期 print("\n===== 评论标题与日期 =====") review_items = soup.find_all('div', class_='review-article') for item in review_items: title = item.find('h2').text.strip() if item.find('h2') else '无标题' date = item.find('span', class_='review-date').text.strip() if item.find('span', class_='review-date') else '无日期' print(f"标题:{title}\n日期:{date}\n") else: print(f"请求失败,状态码:{response.status_code}")
额外提醒
- 如果仍无法获取数据,说明网站用了JavaScript动态渲染内容,需改用
selenium或playwright模拟浏览器加载页面 - 不要高频请求,可加
time.sleep(2)控制间隔,避免IP被封禁 - 遵守网站爬取规则,不要过度抓取数据
内容的提问来源于stack exchange,提问作者dooby
相关产品推荐
相关产品推荐

