使用Rvest爬取Steam评论的两类技术问题求助
Hey there, let's tackle your two scraping issues step by step—they're super common when dealing with dynamic sites like Steam, so you're in good company!
1. Separating Review Content from Post Date
Looking at your code, the review variable pulls the entire text from div.apphub_CardTextContent, which includes both the post date (the small top line) and the actual review body. Since you're already extracting the date separately, we can just strip that date string from the raw review text to get clean content.
Here's how to adjust your code with the stringr package to handle this:
library(rvest) library(stringr) link <- "https://steamcommunity.com/app/215080/reviews/" page <- read_html(link) # Only read the page once for efficiency! # Extract all fields date <- page %>% html_nodes("div.date_posted") %>% html_text() %>% str_trim() # Clean up extra whitespace raw_review <- page %>% html_nodes("div.apphub_CardTextContent") %>% html_text() # Remove the matching date from each review and trim leftover whitespace review <- map2_chr(raw_review, date, ~str_remove_all(.x, .y) %>% str_trim()) opinion <- page %>% html_nodes("div.title") %>% html_text() %>% str_trim() hoursplayed <- page %>% html_nodes("div.hours") %>% html_text() %>% str_trim() helpful <- page %>% html_nodes("div.found_helpful") %>% html_text() %>% str_trim() # Build your data frame tab <- data.frame( "Posted" = date, "Review" = review, "Opinion" = opinion, "Hours Played" = hoursplayed, "Number of helpful vote" = helpful )
A quick note: I adjusted your code to read the page only once instead of multiple times—this saves bandwidth and makes your script faster.
2. Getting All Reviews (Handling Infinite Scroll)
Rvest can only parse static HTML, so it won't pick up content loaded via JavaScript when you scroll. You have two solid options here:
Option 1: Use Steam's Hidden API (Faster & More Reliable)
Steam loads reviews via an AJAX API endpoint, which we can call directly with httr to fetch batches of reviews. This avoids needing to simulate a browser.
Here's a script that loops through all pages of reviews:
library(rvest) library(httr) library(stringr) library(dplyr) app_id <- "215080" base_url <- "https://steamcommunity.com/app/%s/homecontent/" # Initialize storage for all reviews all_reviews <- list() offset <- 0 batch_size <- 10 # Steam returns 10 reviews per request by default while(TRUE) { # Build the API request with pagination parameters url <- sprintf(base_url, app_id) params <- list( userreviewsoffset = offset, numperpage = batch_size, filter = "all", sort = "mostrecent", # Change to "mosthelpful" if needed language = "english", # Adjust to your target language reviewType = "all" ) # Send the request response <- GET(url, query = params) if (http_status(response)$category != "Success") break # Parse the returned HTML page <- read_html(content(response, "text")) # Extract data for this batch date <- page %>% html_nodes("div.date_posted") %>% html_text() %>% str_trim() raw_review <- page %>% html_nodes("div.apphub_CardTextContent") %>% html_text() review <- map2_chr(raw_review, date, ~str_remove_all(.x, .y) %>% str_trim()) opinion <- page %>% html_nodes("div.title") %>% html_text() %>% str_trim() hoursplayed <- page %>% html_nodes("div.hours") %>% html_text() %>% str_trim() helpful <- page %>% html_nodes("div.found_helpful") %>% html_text() %>% str_trim() # Stop if we hit an empty page if (length(date) == 0) break # Add this batch to our list all_reviews[[offset/batch_size + 1]] <- data.frame( Posted = date, Review = review, Opinion = opinion, Hours_Played = hoursplayed, Helpful_Votes = helpful ) # Move to the next batch offset <- offset + batch_size Sys.sleep(2) # Add a small delay to avoid getting blocked by Steam } # Combine all batches into one data frame final_reviews <- bind_rows(all_reviews)
Option 2: Simulate Browser Scrolling with RSelenium
If you prefer to mimic human scrolling (e.g., if the API parameters change), you can use RSelenium to control a browser:
library(RSelenium) library(rvest) library(stringr) # Start a Chrome driver (you'll need ChromeDriver installed first) driver <- rsDriver(browser = "chrome", port = 4567L) remote_driver <- driver[["client"]] # Navigate to the reviews page remote_driver$navigate("https://steamcommunity.com/app/215080/reviews/") # Scroll to load all reviews last_height <- remote_driver$executeScript("return document.body.scrollHeight;") while(TRUE) { remote_driver$executeScript("window.scrollTo(0, document.body.scrollHeight);") Sys.sleep(3) # Wait for new content to load new_height <- remote_driver$executeScript("return document.body.scrollHeight;") if (new_height == last_height) break # Stop when no new content loads last_height <- new_height } # Extract the full page HTML page <- read_html(remote_driver$getPageSource()[[1]]) # Now extract your data using the same method as before date <- page %>% html_nodes("div.date_posted") %>% html_text() %>% str_trim() raw_review <- page %>% html_nodes("div.apphub_CardTextContent") %>% html_text() review <- map2_chr(raw_review, date, ~str_remove_all(.x, .y) %>% str_trim()) # ... extract other fields here # Clean up remote_driver$close() driver$server$stop()
This method is slower but works if the API approach isn't feasible.
内容的提问来源于stack exchange,提问作者KudoJr

