R语言网页爬取标题遇问题:rvest工具无效,寻求解决方案
Hey there! Let's break down why your rvest code is spitting out empty strings when trying to grab the title for Her (2013), and whether switching to RCurl will fix things.
First: Let's Fix Your rvest Setup (RCurl Isn't the Silver Bullet)
RCurl is a lower-level HTTP request package, but it doesn't handle HTML parsing on its own—you'd still need to pair it with something like XML to extract nodes. The issue you're facing is almost certainly not the package itself, but either your node selector, missing request headers, or dynamic content loading. Let's tackle these step by step:
1. Double-Check Your Node Selector
It's super easy to get the wrong CSS/XPath selector. Open your target movie page in a browser, right-click the title, and use "Inspect" to find the exact element:
- If the title is in an
<h1>tag with a class likemovie-title, your CSS selector should beh1.movie-title - For XPath, it might look like
//h1[@class='movie-title']
Test this with a proper request that mimics a browser (many sites block scrapers without a user agent):
library(rvest) library(httr) # Replace with your actual movie page URL movie_url <- "https://your-movie-site.com/her-2013" # Send a request with a browser-like user agent page_request <- GET( movie_url, user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") ) # Parse the response and extract the title movie_page <- read_html(page_request) movie_title <- movie_page %>% html_element("h1.movie-title") %>% # Replace with your correct selector html_text2() # html_text2() handles whitespace better than html_text() print(movie_title)
2. Check If the Content Is Dynamically Loaded
If the above still returns an empty string, the title might be loaded via JavaScript (rvest only parses static HTML). In this case, you'll need to simulate a browser to render the JS:
library(RSelenium) # Start a Chrome driver (make sure ChromeDriver is installed and matches your Chrome version) driver <- rsDriver(browser = "chrome", chromever = "114.0.5735.90") browser <- driver[["client"]] # Navigate to the page and wait for JS to load browser$navigate(movie_url) Sys.sleep(3) # Adjust wait time based on how fast the page loads # Grab the rendered page source and parse it rendered_source <- browser$getPageSource()[[1]] movie_page <- read_html(rendered_source) # Extract the title as before movie_title <- movie_page %>% html_element("h1.movie-title") %>% html_text2() print(movie_title) # Clean up browser$close() driver$server$stop()
What About RCurl?
If you still want to try RCurl, here's how you'd use it (but remember, it won't solve dynamic content issues):
library(RCurl) library(XML) # Set request headers to mimic a browser headers <- c( "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" ) # Fetch the HTML content html_content <- getURL(movie_url, httpheader = headers) # Parse and extract the title with XPath doc <- htmlParse(html_content) movie_title <- xpathSApply(doc, "//h1[@class='movie-title']", xmlValue) print(movie_title)
Quick Debugging Tip
Before diving into parsing, check if your request is even returning the right content:
# For rvest/httr cat(content(page_request, "text")) # For RCurl cat(html_content)
Search the output for the movie title—if it's not there, the issue is with your request (headers, blocked IP, etc.) or dynamic loading, not your selector.
内容的提问来源于stack exchange,提问作者JellisHeRo

