You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言网页爬取标题遇问题:rvest工具无效,寻求解决方案

Troubleshooting Empty String Output When Scraping Movie Titles with rvest (and RCurl Considerations)

Hey there! Let's break down why your rvest code is spitting out empty strings when trying to grab the title for Her (2013), and whether switching to RCurl will fix things.

First: Let's Fix Your rvest Setup (RCurl Isn't the Silver Bullet)

RCurl is a lower-level HTTP request package, but it doesn't handle HTML parsing on its own—you'd still need to pair it with something like XML to extract nodes. The issue you're facing is almost certainly not the package itself, but either your node selector, missing request headers, or dynamic content loading. Let's tackle these step by step:

1. Double-Check Your Node Selector

It's super easy to get the wrong CSS/XPath selector. Open your target movie page in a browser, right-click the title, and use "Inspect" to find the exact element:

  • If the title is in an <h1> tag with a class like movie-title, your CSS selector should be h1.movie-title
  • For XPath, it might look like //h1[@class='movie-title']

Test this with a proper request that mimics a browser (many sites block scrapers without a user agent):

library(rvest)
library(httr)

# Replace with your actual movie page URL
movie_url <- "https://your-movie-site.com/her-2013"

# Send a request with a browser-like user agent
page_request <- GET(
  movie_url,
  user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
)

# Parse the response and extract the title
movie_page <- read_html(page_request)
movie_title <- movie_page %>% 
  html_element("h1.movie-title") %>% # Replace with your correct selector
  html_text2() # html_text2() handles whitespace better than html_text()

print(movie_title)

2. Check If the Content Is Dynamically Loaded

If the above still returns an empty string, the title might be loaded via JavaScript (rvest only parses static HTML). In this case, you'll need to simulate a browser to render the JS:

library(RSelenium)

# Start a Chrome driver (make sure ChromeDriver is installed and matches your Chrome version)
driver <- rsDriver(browser = "chrome", chromever = "114.0.5735.90")
browser <- driver[["client"]]

# Navigate to the page and wait for JS to load
browser$navigate(movie_url)
Sys.sleep(3) # Adjust wait time based on how fast the page loads

# Grab the rendered page source and parse it
rendered_source <- browser$getPageSource()[[1]]
movie_page <- read_html(rendered_source)

# Extract the title as before
movie_title <- movie_page %>% 
  html_element("h1.movie-title") %>% 
  html_text2()

print(movie_title)

# Clean up
browser$close()
driver$server$stop()

What About RCurl?

If you still want to try RCurl, here's how you'd use it (but remember, it won't solve dynamic content issues):

library(RCurl)
library(XML)

# Set request headers to mimic a browser
headers <- c(
  "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
)

# Fetch the HTML content
html_content <- getURL(movie_url, httpheader = headers)

# Parse and extract the title with XPath
doc <- htmlParse(html_content)
movie_title <- xpathSApply(doc, "//h1[@class='movie-title']", xmlValue)

print(movie_title)

Quick Debugging Tip

Before diving into parsing, check if your request is even returning the right content:

# For rvest/httr
cat(content(page_request, "text"))

# For RCurl
cat(html_content)

Search the output for the movie title—if it's not there, the issue is with your request (headers, blocked IP, etc.) or dynamic loading, not your selector.

内容的提问来源于stack exchange,提问作者JellisHeRo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:44:02