You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium抓取烂番茄audience score观众评分的方法

Selenium提取烂番茄影片观众评分实现方法

问题说明

  • 目标:爬取烂番茄(Rotten Tomatoes)站点的影片观众评分(audience score),目前已可正常爬取评论内容,但无法提取目标audiencescore属性值
  • 已知页面规则:观众评分存储在自定义<score-board>标签的audiencescore属性中,对应HTML结构片段如下:
<score-board
audiencestate="upright"
audiencescore="96"
class="scoreboard"
rating="R"
skeleton="panel"
tomatometerstate="certified-fresh"
tomatometerscore="92"
data-qa="score-panel"
                >
<h1 slot="title" class="scoreboard__title" data-qa="score-panel-movie-title">Pulp Fiction</h1>
<p slot="info" class="scoreboard__info">1994, Crime/Drama, 2h 33m</p>
<a slot="critics-count" href="/m/pulp_fiction/reviews?intcmp=rt-scorecard_tomatometer-reviews" class="scoreboard__link scoreboard__link--tomatometer" data-qa="tomatometer-review-count">110 Reviews</a>
<a slot="audience-count" href="/m/pulp_fiction/reviews?type=user&amp;intcmp=rt-scorecard_audience-score-reviews" class="scoreboard__link scoreboard__link--audience" data-qa="audience-rating-count">250,000+ Ratings</a>
<div slot="sponsorship" id="tomatometer_sponsorship_ad"></div>
                </score-board>
  • 现有代码问题:当前编写的代码仅能获取观众评分的参评人数文本,无法直接提取评分数值,现有代码如下:
from selenium import webdriver

driver = webdriver.Firefox()
url = 'https://www.rottentomatoes.com/m/pulp_fiction'
driver.get(url)

print(driver.find_element_by_css_selector('a[slot=audience-count]').text)

解决方法

原有代码拿不到评分的核心原因是定位元素错误:audiencescore属性并不在当前选中的slot=audience-count的<a>标签上,而是在外层的自定义<score-board>标签上,通过Selenium内置的get_attribute()方法即可直接提取属性值。
修正后的可运行代码如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
import time

driver = webdriver.Firefox()
url = 'https://www.rottentomatoes.com/m/pulp_fiction'
driver.get(url)
# 等待页面动态渲染完成,避免元素未加载导致取值失败
time.sleep(3)

# 定位score-board元素,两种定位方式可选,优先选第二种稳定性更高
# score_board = driver.find_element(By.TAG_NAME, 'score-board')
score_board = driver.find_element(By.CSS_SELECTOR, 'score-board[data-qa="score-panel"]')

# 提取目标属性值
audience_score = score_board.get_attribute('audiencescore')
# 可同步提取烂番茄专业评分
tomatometer_score = score_board.get_attribute('tomatometerscore')
# 原有逻辑提取参评人数
audience_count = driver.find_element(By.CSS_SELECTOR, 'a[slot="audience-count"]').text

print(f"观众评分:{audience_score}")
print(f"专业评分(烂番茄新鲜度):{tomatometer_score}")
print(f"参评规模:{audience_count}")

driver.quit()

注意事项

  • Selenium 4及以上版本已经废弃了find_element_by_css_selector、find_element_by_tag_name这类旧写法,统一使用find_element(定位方式, 定位表达式)的语法,避免版本兼容报错
  • 烂番茄页面为前端动态渲染,页面加载完成后才会给score-board挂载对应属性,建议加等待逻辑,不要页面刚发起请求就立刻取元素值

内容的提问来源于stack exchange,提问作者Lacer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 13:12:17