使用BeautifulSoup爬取Bed Bath & Beyond评论返回空值求助
解决Bed Bath & Beyond评论爬取空列表的问题
嘿,我来帮你搞定这个爬虫卡壳的问题!你遇到的情况很典型——电商网站的评论几乎都是动态加载的,不是直接写在初始页面的HTML里,所以直接用BeautifulSoup解析静态页面,甚至用PyQt4渲染时没等异步请求完成,自然拿不到数据。下面给你具体的解决思路和代码示例:
核心原因:评论是AJAX异步加载的
当你打开商品页面时,初始HTML只包含商品基本信息(描述、价格这些),评论数据是页面加载完成后,浏览器通过单独的AJAX请求从后端API接口拉取的。所以你直接解析初始页面,肯定找不到评论内容。
方案1:直接请求评论API接口(推荐,高效稳定)
这是最靠谱的方式,直接跳过页面渲染,去抓后端给前端喂数据的接口:
- 抓包找API:打开浏览器F12切换到「Network」标签,刷新页面后,筛选「XHR」类型的请求,找带「review」「comment」或产品ID(1061083288)的请求。你会发现一个返回JSON格式数据的接口,这就是评论的数据源。
- 用Python请求API:用
requests库模拟浏览器请求这个接口,就能拿到完整的评论数据。示例代码如下:
import requests import json import time # 替换成你抓包找到的真实评论API地址 review_api = "https://www.bedbathandbeyond.com/api/product/reviews?productId=1061083288&page=1&limit=20" # 模拟浏览器请求头,避免被识别为爬虫 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/90.0.4430.93 Safari/537.36", "Accept": "application/json, text/plain, */*" } try: response = requests.get(review_api, headers=headers) response.raise_for_status() # 检查请求是否成功 review_data = json.loads(response.text) # 提取评论和评论者位置 for item in review_data.get("reviews", []): print("评论内容:", item.get("comment")) print("评论者位置:", item.get("reviewerLocation")) print("---") # 如果有多页评论,循环请求即可 # total_pages = review_data.get("totalPages") # for page in range(2, total_pages+1): # time.sleep(2) # 加延迟,避免被封 # page_api = f"{review_api}&page={page}" # # 重复上述请求逻辑 except Exception as e: print(f"请求出错:{str(e)}")
方案2:优化PyQt4页面渲染(适合坚持用渲染的场景)
如果你还是想用PyQt4渲染页面,必须等AJAX请求完成后再抓取HTML。可以通过定时器延迟抓取,或者监听页面元素加载状态:
from PyQt4.QtGui import QApplication from PyQt4.QtCore import QUrl, QTimer from PyQt4.QtWebKit import QWebPage import sys from bs4 import BeautifulSoup class PageRenderer(QWebPage): def __init__(self, target_url): self.app = QApplication(sys.argv) QWebPage.__init__(self) self.load_finished = False self.loadFinished.connect(self.on_load_finished) self.mainFrame().load(QUrl(target_url)) # 给AJAX留足够加载时间,比如5秒 QTimer.singleShot(5000, self.force_quit) self.app.exec_() def on_load_finished(self, result): self.html_content = self.mainFrame().toHtml() self.load_finished = True self.app.quit() def force_quit(self): if not self.load_finished: self.html_content = self.mainFrame().toHtml() self.app.quit() # 渲染页面 target_url = 'https://www.bedbathandbeyond.com/store/product/dyson-v7-motorhead-cord-free-stick-vacuum-in-fuchsia-steel/1061083288?brandId=162' renderer = PageRenderer(target_url) soup = BeautifulSoup(renderer.html_content, 'lxml') # 替换成页面中评论和位置对应的真实CSS选择器 reviews = soup.find_all(class_="review-body-text") review_locations = soup.find_all(class_="reviewer-location") print("评论列表:", [r.get_text(strip=True) for r in reviews]) print("评论者位置:", [l.get_text(strip=True) for l in review_locations])
不过这种方法效率低,而且Python2.7的PyQt4对现代JS的支持有限,还是优先选API方案。
额外提醒
- 别忘加请求延迟(
time.sleep()),避免频繁请求被网站封禁IP。 - 如果API需要Cookie或Token,抓包时要把请求头里的对应字段复制过来。
- 另外,Python2.7已经停止维护了,能升级到Python3.x的话,库的支持会好很多哦!
内容的提问来源于stack exchange,提问作者Eric Cai
相关产品推荐
相关产品推荐

