You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium和BeautifulSoup爬取《饥饿游戏》用户评分遇重复数据问题

解决Goodreads评论爬取时重复第一页数据的问题

我太懂你遇到的这个坑了——用Selenium翻页爬评论,结果输出全是第一页的重复内容,折腾半天白忙活!咱们来拆解问题,快速搞定它。

问题根源

你现在的代码里,只在最开始解析了一次页面的HTML,之后翻页后完全没更新页面数据。循环里一直用第一次获取的user列表(也就是第一页的评论),自然每次循环都把第一页的数据再塞一遍到ratings里,结果全是重复内容。另外还有两个小隐患:没提前初始化ratings列表,以及翻页后按钮元素可能失效(变成stale element)。

修正后的代码

下面是调整后的完整代码,我加上了详细注释,解决了所有问题:

from selenium import webdriver
from bs4 import BeautifulSoup
import pandas as pd
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 初始化Chrome驱动
path_to_chromedriver = r'./chromedriver.exe'
url = "https://www.goodreads.com/book/show/2767052-the-hunger-games"
driver = webdriver.Chrome(executable_path=path_to_chromedriver)
driver.implicitly_wait(30)
driver.get(url)

# 提前初始化存储评分的列表,避免未定义错误
ratings = []

# 假设总共有10页评论,循环10次(包含第一页到第十页)
for page_num in range(10):  
    # 关键:每次循环都重新解析当前页面的源码!
    soup = BeautifulSoup(driver.page_source, 'lxml')
    reviews_container = soup.find('div', {'id': 'bookReviews'})
    # 获取当前页的所有评论元素
    user_reviews = reviews_container.find_all('div', {'class':'friendReviews elementListBrown'})
    
    # 遍历当前页的评论,提取数据
    for row in user_reviews:
        rating = {}
        try:
            rating['name'] = row.find('a', {'class': 'user'}).text.strip()
            rating_title = row.find('span', {'class':'staticStars'})['title']
            rating['rating'] = rating_title
            # 如果你只需要五星评分,加上这个判断(Goodreads五星对应"it was amazing")
            if rating_title == "it was amazing":
                ratings.append(rating)
        except Exception as e:
            # 打印错误方便排查,也可以直接pass
            print(f"跳过一条评论,原因:{str(e)}")
            pass
    
    # 不是最后一页的话,点击下一页
    if page_num != 9:
        try:
            # 用显式等待确保按钮可点击,解决翻页后元素失效的问题
            next_btn = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.PARTIAL_LINK_TEXT, 'next »'))
            )
            next_btn.click()
            # 短暂等待页面加载(也可以用显式等待评论容器重新出现,更严谨)
            driver.implicitly_wait(5)
        except Exception as e:
            print(f"点击下一页失败,原因:{str(e)}")
            break

# 转换为DataFrame并查看结果
df_rev = pd.DataFrame(ratings)
print(df_rev)
# 记得关闭浏览器释放资源
driver.quit()

关键修改点说明

  1. 每次循环重新解析页面:把BeautifulSoup解析代码放到循环内部,确保每次处理的都是当前页的最新评论数据。
  2. 初始化评分列表:提前定义ratings = [],避免潜在的未定义变量错误。
  3. 显式等待定位下一页按钮:用WebDriverWait代替直接查找元素,解决翻页后按钮元素失效的问题,爬取更稳定。
  4. 可选的五星过滤逻辑:因为你的目标是获取五星评分,加了判断只保留对应title的评论,不需要的话可以删掉这个if。

内容的提问来源于stack exchange,提问作者realkes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:26:46