You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言rvest爬取IMDB尼日利亚电影演员与导演失败求助

解决IMDB多页尼日利亚电影演员与导演数据爬取问题

核心问题排查

你遇到的「仅首个电影有数据、导演爬取返回演员数据」问题,通常源于两个原因:

  • 循环中未独立解析每个电影的fullcredits页面,导致选择器复用了第一个页面的DOM结构
  • 演员/导演的CSS选择器混淆,爬导演时误用了演员的选择器

修正后的完整爬取代码

使用rvest+purrr实现稳定的多页爬取,同时解决反爬、选择器错误问题:

library(rvest)
library(dplyr)
library(purrr)
library(stringr)
library(glue)

# 设置请求头,模拟浏览器访问
headers <- c(
  "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
)

# 定义单电影演员/导演提取函数
get_crew_cast <- function(fullcredits_url) {
  # 添加延迟,规避IMDB反爬机制
  Sys.sleep(1)
  
  # 独立解析当前电影的fullcredits页面
  page <- read_html(fullcredits_url, headers = headers)
  
  # 提取主要导演:定位Director标题后的表格内的姓名链接
  directors <- page %>%
    html_elements("h4:contains('Director') + table .name a") %>%
    html_text2() %>%
    paste(collapse = ", ")
  
  # 提取所有演员:定位Cast标题后的表格内的姓名链接
  cast <- page %>%
    html_elements("h4:contains('Cast') + table .name a") %>%
    html_text2() %>%
    paste(collapse = ", ")
  
  return(tibble(directors = directors, cast = cast))
}

# 爬取多页数据(示例为前3页,可自行扩展start参数)
start_pages <- c(1, 51, 101)

all_movies_data <- map_dfr(start_pages, function(start) {
  # 构造当前页的主链接
  main_url <- glue("https://www.imdb.com/search/title/?country_of_origin=NG&start={start}&ref_=adv_prv")
  main_page <- read_html(main_url, headers = headers)
  
  # 提取基础信息
  base_info <- tibble(
    title = main_page %>% html_elements(".lister-item-header a") %>% html_text2(),
    year = main_page %>% html_elements(".lister-item-year") %>% html_text2() %>% str_extract("\\d{4}"),
    summary = main_page %>% html_elements(".ratings-bar + .text-muted") %>% html_text2() %>% str_trim(),
    genre = main_page %>% html_elements(".genre") %>% html_text2() %>% str_trim(),
    rating = main_page %>% html_elements(".certificate") %>% html_text2(),
    # 构造fullcredits链接
    fullcredits_url = main_page %>% html_elements(".lister-item-header a") %>% html_attr("href") %>%
      paste0("https://www.imdb.com", ., "/fullcredits/?ref_=tt_cl_sm")
  )
  
  # 批量获取每个电影的演员/导演数据
  crew_cast_data <- map_dfr(base_info$fullcredits_url, get_crew_cast)
  
  # 合并基础信息与演员导演数据
  bind_cols(base_info, crew_cast_data) %>% select(-fullcredits_url)
})

# 查看结果
print(all_movies_data)

关键修正说明

  1. 独立页面解析:每个电影的fullcredits页面都单独调用read_html,避免DOM结构污染
  2. 精准选择器:
    • 导演选择器h4:contains('Director') + table .name a:定位到「Director」标题后的表格内的姓名链接
    • 演员选择器h4:contains('Cast') + table .name a:定位到「Cast」标题后的表格内的姓名链接
  3. 反爬处理:添加Sys.sleep(1)延迟,避免请求频率过高被IMDB拦截
  4. 批量处理:用map_dfr替代手动循环,避免索引错误,同时自动合并数据框

调试建议

  • 若仍有部分电影无数据,单独打印该电影的fullcredits_url,检查页面结构是否特殊(如无Director标签),可添加ifelse逻辑处理空值
  • 若请求被拦截,可延长Sys.sleep的时间,或更换User-Agent

内容的提问来源于stack exchange,提问作者Emmanuel Ogebe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 00:42:13