You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Bed Bath & Beyond评论返回空值求助

解决Bed Bath & Beyond评论爬取空列表的问题

嘿,我来帮你搞定这个爬虫卡壳的问题!你遇到的情况很典型——电商网站的评论几乎都是动态加载的,不是直接写在初始页面的HTML里,所以直接用BeautifulSoup解析静态页面,甚至用PyQt4渲染时没等异步请求完成,自然拿不到数据。下面给你具体的解决思路和代码示例:

核心原因:评论是AJAX异步加载的

当你打开商品页面时,初始HTML只包含商品基本信息(描述、价格这些),评论数据是页面加载完成后,浏览器通过单独的AJAX请求从后端API接口拉取的。所以你直接解析初始页面,肯定找不到评论内容。

方案1:直接请求评论API接口(推荐,高效稳定)

这是最靠谱的方式,直接跳过页面渲染,去抓后端给前端喂数据的接口:

  1. 抓包找API:打开浏览器F12切换到「Network」标签,刷新页面后,筛选「XHR」类型的请求,找带「review」「comment」或产品ID(1061083288)的请求。你会发现一个返回JSON格式数据的接口,这就是评论的数据源。
  2. 用Python请求API:用requests库模拟浏览器请求这个接口,就能拿到完整的评论数据。示例代码如下:
import requests
import json
import time

# 替换成你抓包找到的真实评论API地址
review_api = "https://www.bedbathandbeyond.com/api/product/reviews?productId=1061083288&page=1&limit=20"

# 模拟浏览器请求头,避免被识别为爬虫
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/90.0.4430.93 Safari/537.36",
    "Accept": "application/json, text/plain, */*"
}

try:
    response = requests.get(review_api, headers=headers)
    response.raise_for_status()  # 检查请求是否成功
    review_data = json.loads(response.text)
    
    # 提取评论和评论者位置
    for item in review_data.get("reviews", []):
        print("评论内容:", item.get("comment"))
        print("评论者位置:", item.get("reviewerLocation"))
        print("---")
        
    # 如果有多页评论,循环请求即可
    # total_pages = review_data.get("totalPages")
    # for page in range(2, total_pages+1):
    #     time.sleep(2)  # 加延迟,避免被封
    #     page_api = f"{review_api}&page={page}"
    #     # 重复上述请求逻辑
except Exception as e:
    print(f"请求出错:{str(e)}")

方案2:优化PyQt4页面渲染(适合坚持用渲染的场景)

如果你还是想用PyQt4渲染页面,必须等AJAX请求完成后再抓取HTML。可以通过定时器延迟抓取,或者监听页面元素加载状态:

from PyQt4.QtGui import QApplication
from PyQt4.QtCore import QUrl, QTimer
from PyQt4.QtWebKit import QWebPage
import sys
from bs4 import BeautifulSoup

class PageRenderer(QWebPage):
    def __init__(self, target_url):
        self.app = QApplication(sys.argv)
        QWebPage.__init__(self)
        self.load_finished = False
        self.loadFinished.connect(self.on_load_finished)
        self.mainFrame().load(QUrl(target_url))
        # 给AJAX留足够加载时间,比如5秒
        QTimer.singleShot(5000, self.force_quit)
        self.app.exec_()

    def on_load_finished(self, result):
        self.html_content = self.mainFrame().toHtml()
        self.load_finished = True
        self.app.quit()

    def force_quit(self):
        if not self.load_finished:
            self.html_content = self.mainFrame().toHtml()
            self.app.quit()

# 渲染页面
target_url = 'https://www.bedbathandbeyond.com/store/product/dyson-v7-motorhead-cord-free-stick-vacuum-in-fuchsia-steel/1061083288?brandId=162'
renderer = PageRenderer(target_url)
soup = BeautifulSoup(renderer.html_content, 'lxml')

# 替换成页面中评论和位置对应的真实CSS选择器
reviews = soup.find_all(class_="review-body-text")
review_locations = soup.find_all(class_="reviewer-location")

print("评论列表:", [r.get_text(strip=True) for r in reviews])
print("评论者位置:", [l.get_text(strip=True) for l in review_locations])

不过这种方法效率低,而且Python2.7的PyQt4对现代JS的支持有限,还是优先选API方案。

额外提醒

  • 别忘加请求延迟(time.sleep()),避免频繁请求被网站封禁IP。
  • 如果API需要Cookie或Token,抓包时要把请求头里的对应字段复制过来。
  • 另外,Python2.7已经停止维护了,能升级到Python3.x的话,库的支持会好很多哦!

内容的提问来源于stack exchange,提问作者Eric Cai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:55:09