You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用rvest从IMDB电影页面抓取影片类型?

解决IMDB影片类型抓取问题

原代码中抓取类型的逻辑存在问题,IMDB当前页面里,影片类型在data-testid="storyline-genres"节点下是以多个独立的链接元素存在的,而非单个div的拼接文本。你可以通过以下方式修正:

修正类型抓取逻辑

替换原代码中movie_genres的赋值部分为:

movie_genres <- movie_page %>%
  html_elements('[data-testid="storyline-genres"] a') %>%
  html_text() %>%
  paste(collapse = ", ")

这段代码的作用:

  • 定位到类型区域下的所有链接元素(每个链接对应一个影片类型)
  • 提取每个类型的文本内容
  • 将多个类型用逗号拼接成统一字符串

完整修正后的函数

将原函数中相关部分替换后,完整代码如下:

# Load necessary libraries
library(rvest)
library(dplyr)

# Function to scrape movie information from IMDB
scrape_imdb <- function(movie_names) {
  base_url <- "https://www.imdb.com/find?q="
  
  # Initialize an empty data frame to store the results
  results_df <- data.frame(Movie = character(),
                           Year = character(),
                           Rating = character(),
                           Genres = character(),
                           Runtime = character(),
                           stringsAsFactors = FALSE)
  
  for (movie_name in movie_names) {
    # Construct search URL
    search_url <- paste0(base_url, URLencode(movie_name), "&s=tt")
    
    # Read the search result page
    search_page <- read_html(search_url)
    
    # Extract the first movie result's URL
    movie_url <- search_page %>%
      html_nodes(".ipc-metadata-list-summary-item__c a") %>%
      html_attr("href") %>%
      .[1] %>%
      paste0("https://www.imdb.com", .)
    
    movie_year <- search_page %>%
      html_nodes(".ipc-metadata-list-summary-item__c span")  %>% 
      html_text() %>%
      .[1]
    
    # Read the movie page
    movie_page <- read_html(movie_url)
    
    movie_rating <- movie_page %>%
      html_nodes(".sc-bde20123-1.cMEQkK") %>%
      html_text() %>%
      .[1]
    
    movie_runtime  <- movie_page %>%
      html_nodes(".sc-d8941411-2.cdJsTz li") %>%
      html_text() %>% 
      .[3]
    
    # 修正后的类型抓取逻辑
    movie_genres <- movie_page %>%
      html_elements('[data-testid="storyline-genres"] a') %>%
      html_text() %>%
      paste(collapse = ", ")
    
    print(movie_genres)
    
    # Append the results to the data frame
    results_df <- rbind(results_df, data.frame(Movie = movie_name,
                                               Year = as.numeric(movie_year),
                                               Rating = as.numeric(movie_rating),
                                               Genres = movie_genres,
                                               Runtime = movie_runtime,
                                               stringsAsFactors = FALSE))
  }
  
  return(results_df)
}

# Example usage
movie_names <- c("The Matrix", "Inception", "Interstellar")

movie_info <- scrape_imdb(movie_names)

print(movie_info)

额外稳定性优化提示

IMDB的动态类名(如.sc-bde20123-1.cMEQkK)会随时变化,建议优先使用data-testid属性定位元素,比如评分可以改用更稳定的选择器:

movie_rating <- movie_page %>%
  html_element('[data-testid="hero-rating-bar__aggregate-rating__score"]') %>%
  html_text() %>%
  stringr::str_extract("\\d+\\.*\\d*") # 提取纯数字评分部分

内容的提问来源于stack exchange,提问作者zaneywolf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 12:27:47