如何用rvest从IMDB电影页面抓取影片类型?
解决IMDB影片类型抓取问题
原代码中抓取类型的逻辑存在问题,IMDB当前页面里,影片类型在data-testid="storyline-genres"节点下是以多个独立的链接元素存在的,而非单个div的拼接文本。你可以通过以下方式修正:
修正类型抓取逻辑
替换原代码中movie_genres的赋值部分为:
movie_genres <- movie_page %>% html_elements('[data-testid="storyline-genres"] a') %>% html_text() %>% paste(collapse = ", ")
这段代码的作用:
- 定位到类型区域下的所有链接元素(每个链接对应一个影片类型)
- 提取每个类型的文本内容
- 将多个类型用逗号拼接成统一字符串
完整修正后的函数
将原函数中相关部分替换后,完整代码如下:
# Load necessary libraries library(rvest) library(dplyr) # Function to scrape movie information from IMDB scrape_imdb <- function(movie_names) { base_url <- "https://www.imdb.com/find?q=" # Initialize an empty data frame to store the results results_df <- data.frame(Movie = character(), Year = character(), Rating = character(), Genres = character(), Runtime = character(), stringsAsFactors = FALSE) for (movie_name in movie_names) { # Construct search URL search_url <- paste0(base_url, URLencode(movie_name), "&s=tt") # Read the search result page search_page <- read_html(search_url) # Extract the first movie result's URL movie_url <- search_page %>% html_nodes(".ipc-metadata-list-summary-item__c a") %>% html_attr("href") %>% .[1] %>% paste0("https://www.imdb.com", .) movie_year <- search_page %>% html_nodes(".ipc-metadata-list-summary-item__c span") %>% html_text() %>% .[1] # Read the movie page movie_page <- read_html(movie_url) movie_rating <- movie_page %>% html_nodes(".sc-bde20123-1.cMEQkK") %>% html_text() %>% .[1] movie_runtime <- movie_page %>% html_nodes(".sc-d8941411-2.cdJsTz li") %>% html_text() %>% .[3] # 修正后的类型抓取逻辑 movie_genres <- movie_page %>% html_elements('[data-testid="storyline-genres"] a') %>% html_text() %>% paste(collapse = ", ") print(movie_genres) # Append the results to the data frame results_df <- rbind(results_df, data.frame(Movie = movie_name, Year = as.numeric(movie_year), Rating = as.numeric(movie_rating), Genres = movie_genres, Runtime = movie_runtime, stringsAsFactors = FALSE)) } return(results_df) } # Example usage movie_names <- c("The Matrix", "Inception", "Interstellar") movie_info <- scrape_imdb(movie_names) print(movie_info)
额外稳定性优化提示
IMDB的动态类名(如.sc-bde20123-1.cMEQkK)会随时变化,建议优先使用data-testid属性定位元素,比如评分可以改用更稳定的选择器:
movie_rating <- movie_page %>% html_element('[data-testid="hero-rating-bar__aggregate-rating__score"]') %>% html_text() %>% stringr::str_extract("\\d+\\.*\\d*") # 提取纯数字评分部分
内容的提问来源于stack exchange,提问作者zaneywolf
相关产品推荐
相关产品推荐

